European Journal of Human Genetics
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match European Journal of Human Genetics's content profile, based on 58 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
MERCIER, S.; PETIT, F.; MISRAHI, M.; BERTA, P.; CAMBON-THOMSEN, A.; CHAUMETTE, B.; CHNEIWEISS, H.; CRETOLLE, C.; EDERY, P.; HEARD, D.; KONYUKH, M.; LAENG, C.; MAHLAOUI, N.; PASQUIER, L.; PLUTINO, M.; ODENT, S.; STOPPA-LYONNET, D.; "Genetics and the General Public" FFGH Ethics Working Group,
Show abstract
Advances in high-throughput sequencing and genetic research have expanded the role of genetics in medicine and society. Population-based screening programs, including neonatal and preconception testing, are increasingly implemented globally, alongside the rise of direct-to-consumer (DTC) genetic testing. The "Genetics and the General Public" Ethics Working Group of the French Federation of Human Genetics (FFGH) assessed knowledge and awareness of genetics within the French population through a nationally representative survey (n=3,013) conducted by the polling firm Ipsos bva. Results indicated that 69% of respondents report an interest in genetics, although their level of knowledge remains limited. Most respondents expressed positive attitudes toward genetics, perceiving it as a major source of hope in healthcare. While a majority indicated willingness to undergo genetic testing for medical purposes, they also reported legitimate concerns regarding the potential results. Despite legal restrictions, 12% reported having ordered a DTC genetic test (5% for genealogical; 5% for medical and 2% for both purposes), and 45% of non-users expressed strong interest in this type of test. Notably, there is a substantial lack of awareness regarding the limitations of these tests and the French legal framework governing their use. These findings highlight critical gaps in public knowledge, emphasizing the need for improved genetic education, including incorporating genetics into school curricula and launching targeted awareness campaigns. These initiatives should help clarify the distinctions between clinically validated genetic tests and DTC genetic testing services, addressing both their benefits and their ethical, legal, and scientific limitations, in order to promote informed decision-making.
Pedersen, E. M.; Steinbach, J.; Valstad, M.; Ohlsson, H.; Rasmussen, L. A.; Eilertsen, E. M.; Kendler, K. S.; Vilhjalmsson, B. J.; Schork, A. J.; Krebs, M. D.
Show abstract
Estimates of per-individual genetic liability from large-scale family data are routinely used in biomedical research to describe genetic etiology of traits and disorders, boost the power of gene-mapping studies, and improve risk predictions. Here we present LT-FGRS, an R package for handling population-scale pedigrees and implementing multiple state-of-the-field methods for estimating genetic liability from such data. Benchmarking in population-scale Nordic registry data demonstrates that LT-FGRS reproduces estimates from existing implementations at manageable computational cost. LT-FGRS unifies previous parallel implementations into a single framework, lowering barriers for methodological comparison and applied use. Availability and ImplementationLT-FGRS is available as an R-package on CRAN. (https://CRAN.R-project.org/package=LTFGRS) Contactemp@au.dk; morten.dybdahl.krebs@regionh.dk Supplementary informationhttps://emilmip.github.io/LTFGRS
Sümer, A. P.; Iasi, L. N. M.; Bossoms Mesa, A.; Slon, V.; Essel, E.; Hajdinjak, M.; Zorn, J.; Schmidt, A.; Nagel, S.; Nickel, B.; Viola, B.; Ziganshin, R.; Buzhilova, A.; Derevianko, A.; Pääbo, S.; Peter, B. M.
Show abstract
The Teshik-Tash 1 child whose remains were found in Uzbekistan represents the southeastern-most extent of the known Neandertal range, providing an important link with the better studied Caucasus and Altai Mountain ranges. However, due to poor DNA preservation, studying the genetics of Teshik Tash 1 has remained elusive. Here we present analyses of the nuclear DNA from the Teshik-Tash 1, from extracts that are highly contaminated with present-day human DNA. To achieve this, we developed a new computational method, admixslug, that jointly models contamination and population relationships, in order to infer the relationship of a target individual from which only low-quality nuclear DNA is available, to high-quality archaic human genomes. After validating admixslug, we show that Teshik-Tash 1 is genetically more similar to later Neandertals from Western Eurasia than to older Neandertals from the Altai Mountains. We estimate that Teshik-Tash 1 split from the Western Eurasian lineage between 80,000 and 100,000 years ago. Despite the geographical proximity of Teshik-Tash 1 to the Denisovan range, we find no evidence for Denisovan ancestry in his genome. Our results demonstrate that admixslug enables the study of archaic human specimens in cases where DNA preservation was previously considered too poor for population genetic analyses.
Maxwell, G. E.; Allen, R.; Hodge, L.; Kelley, S.; Craig, J. E.; Cohen-Woods, S.; Souzeau, E.
Show abstract
Early glaucoma detection and treatment are critical to prevent irreversible blindness. Glaucoma polygenic risk scores (PRS) offer an effective approach for stratifying disease risk and are increasingly available in clinical practice. However, the psychosocial impact of receiving glaucoma PRS results is currently unknown. As such, this study investigated short-term psychosocial outcomes of disclosing glaucoma PRS to individuals over 50 years from the general population. Individuals from the bottom 10%, middle 45 to 55%, and top 10% of PRS scores were invited to receive their results and complete surveys before and 2 weeks after receiving results to assess anxiety, test-related distress, decisional regret, recall and understanding. Of invited participants, 51.7% (136/263) enrolled with 133 completing both surveys. Two weeks after disclosure, PRS recall was high (78.2%), although PRS knowledge remained limited. Privacy concerns were moderate, not differing across PRS groups (X^2 = 4.17, p = .124). Small reductions in glaucoma-related anxiety (z = -2.93, p = .003), generalised anxiety (z = -3.75, p < .001) and stress (z = -2.49, p = .013) were observed following disclosure. While scores remained within normal ranges, higher glaucoma-related anxiety (t = -2.36, p = .020), higher negative emotions (X^2 = 20.80, p < .001), and lower positive experience (F = 5.70, p = .004) were seen for high-risk participants compared to lower-risk participants. Decisional regret was low and did not differ across PRS groups (X^2 = 0.28, p = .869). These findings support the psychosocial safety of glaucoma PRS testing while highlighting the need for improved education and longer-term follow-up to support clinical implementation.
Adegbesan, A. C.; FitzGerald, L.; Dickinson, J. L.; Raspin, K.; Roydhouse, J.
Show abstract
Background: Patient-reported measures (PRMs), including patient-reported outcome and experience measures, capture patients perspectives on their health status and healthcare experiences. In cancer genetics, PRMs have been used to assess genetic knowledge, psychosocial outcomes, and decision-making. However, patients must understand these measures to provide useful information, an ability which is influenced by general and health literacy levels. Readability guidelines recommend that patient-facing materials be written at or below a Grade 6 level. This study evaluated the readability of PRMs used in a cancer genetic testing context. Objective: To assess whether PRMs used in heritable cancer genetic testing meet recommended readability levels using validated indices. Methods: PRMs were identified from a recent systematic review of PRMs used in heritable cancer genetic testing, which reported 83 instruments across eight categories. English-language PRMs containing structured question items and response scales were eligible for extraction and converted into plain text for analysis. Readability was assessed using four validated indices: Flesch Kincaid Grading Level (FKGL), FORd, CAylor, and STicht (FORCAST) formula, Flesch Reading Ease Score (FRES), and Simple Measure of Gobbledygook (SMOG) via an automated readability software. Descriptive analysis and numerical comparison evaluated readability levels across PRM categories and against the recommended Grade 6 reading level. Results: Sixty-five PRMs met the eligibility criteria, with most, including validated instruments, exceeding the recommended Grade 6 reading level. Across the eight categories, genetics-specific PRMs required the highest readability levels, indicating higher readability demands. Conclusions: Most PRMs, particularly those specific to genetics, do not meet readability guidelines. This may limit their accessibility to individuals with limited general and health literacy. Development of PRMs specific to genetics should consider strategies to improve readability, such as plain-language approaches and involvement of individuals with limited general or health literacy. Keywords: readability, patient-reported measures, cancer, genetic testing, health literacy
Linderman, M. D.; Adelson, S. M.; Berro, T. M.; Anderson, J. L.; Crawford, S. D.; Cunningham, T. J.; Esplin, E. D.; Ewing-Crawford, A. T.; Nielsen, D. E.; Pereira, S.; Schmidlen, T.; Andrighetti, H.; Bleyl, S. B.; Church, G. M.; Haverfield, E. V.; Hegde, M.; Konstantinos, L. N.; Kruszka, P.; Leonard, D.; May, T.; McGinniss, M.; Pandya, V.; Schadt, E. E.; Greshake Tzovaras, B.; Zettler, B.; McGuire, A. L.; Green, R. C.; PeopleSeq Study Team,
Show abstract
Purpose: Elective genomic sequencing (EGS) returns monogenic disease findings in multiple genes, including potentially novel variants, and may also provide participants with carrier status, pharmacogenomic and other health-related information. The PeopleSeq Study assessed participants' motivations for and concerns about EGS and the associated clinical and psychosocial outcomes across diverse EGS providers. Methods: We administered a shared questionnaire to participants who chose to undergo EGS via 18 academic, clinical, or commercial EGS platforms. Results: We enrolled 1575 participants, of whom 1147 (72.8%) completed a questionnaire after receiving their EGS results. A majority (60.3%) of the participants who completed a post-result questionnaire self-reported receiving results they assessed as important, including negative findings, and 75.9% reported a form of health-related utility. Among a subset (19.4%) who shared their EGS reports, 16.6% (37 of n=223) received a monogenic finding and self-reported results deemed "important" were consistent with EGS reports. Most participants (74.1%) discussed their results with their family, but fewer discussed their results with a healthcare provider other than the site team (41.7%) or had one or more medical visits as a direct result of their EGS testing (23.1%). Participants expressed diverse motivations for EGS, with 91.4% expressing interest in their personal disease risk and 54% who expressed quasi-indication-based motivations related to family medical history. Individuals motivated by family history reported important results at a significantly higher rate. Conclusions: Early adopters of EGS are motivated by general interest in their health as well as quasi-indication-based considerations such as family history. A majority of participants learned results they considered medically important, but a much smaller segment engaged healthcare providers with their results.
Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.
Show abstract
Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.
Sangkuhl, K.; Whirl-Carrillo, M.; Woon, M.; Venkatesh, R.; Keat, K.; Whaley, R.; Ritchie, M. D.; Klein, T. E.
Show abstract
NAT2 is an important pharmacogene which encodes the N-acetyltransferase 2 enzyme that is involved in the metabolism of multiple medications, and variants in this gene can affect patient response to these medications. CPIC has published a clinical guideline for prescribing hydralazine using NAT2 genotypes. Just prior to the guideline, updated NAT2 star allele numbering and definitions were released, differing somewhat from the historical nomenclature. Clinical pharmacogenomic testing panels often test for the most common star alleles, so knowledge of the most common updated NAT2 star alleles is critical for the implementation of the CPIC NAT2/hydralazine guideline. We first determine NAT2 diplotype frequencies from UK Biobank (UKBB) 200k phased genomes, then analyzed allele, diplotype, and phenotype population frequencies from the All of Us Research program, PennMedicine BioBank (PMBB) and UKBB 500k datasets. We found that analyzing NAT2 diplotypes from phased data provides critical information for algorithms designed to predict diplotypes from unphased data. We observed that NAT2*5, *6, and *4 were the most common star alleles in that order, and the top 11 most frequent NAT2 star alleles were the same across all biobanks. However, differences in star allele frequencies across biogeographical populations were observed. The largest difference led to a higher frequency of NAT2 poor metabolizer phenotypes as compared to rapid and intermediate metabolizer phenotypes in all global populations except in the EAS population, where NAT2 poor metabolizers were in the minority.
Delagrammatikas, C. G.; Gourlay, L. J.; Priolo, M.; Russo, R.; Ahmadi, A.; Barbiroli, A. G.; Capelli, R.; Stowers, K.; D'Annibale, O.; Ravalin, M.; Tartaglia, M.; Nardini, M.; Cocanougher, B. T.
Show abstract
Purpose: Pathogenic variants in NFIX cause Marshall-Smith syndrome and Malan syndrome (MALNS). We identified a severe subtype of MALNS characterized by adolescent-onset musculoskeletal deterioration and investigated functional consequences of underlying variants. Methods: Clinical data were collected from seven individuals with pathogenic NFIX variants. Wild-type and mutated recombinant NFIX DNA-binding domains (DBDs) were evaluated using biochemical, structural, and DNA-binding assays. Results: Six individuals carrying R116W, R116P, K125E, or G147E NFIX substitutions developed progressive muscle wasting, markedly reduced body mass index, and rapidly progressive scoliosis after the typical childhood features of MALNS; two died from disease-related complications. A seventh individual with R116G did not develop this severe phenotype. Functional studies on recombinant NFIX DBDs showed complete or near-complete loss of DNA-binding activity for R116W, R116P, K125E, and G147E despite preserved protein folding, consistent with disrupted DNA recognition and a potential dominant-negative mechanism. In contrast, R116G exhibited a 7.7{degrees}C decrease in thermal stability, which may support haploinsufficiency mediated by protein degradation. Conclusion: Specific NFIX missense variants define a severe subtype of MALNS associated with progressive musculoskeletal deterioration. In vitro functional studies support variant-specific disruption of DNA binding, providing a mechanistic basis of genotype-phenotype correlations and informing prognosis, clinical surveillance, and therapy development.
Hartman, N. R.; Villanea, F. A.
Show abstract
Adrenarche, the pre-pubertal rise in adrenal androgens, particularly dehydroepiandrosterone (DHEA) and its sulfated form DHEAS, is a critical driver of middle childhood cognitive and social development in modern humans. Compared to other apes, modern human adrenarche is more prolonged, with higher DHEAS levels. Whether this uniquely prolonged human adrenarche is a derived trait of Homo sapiens or has deeper hominin roots remains unresolved. Here, we examine the Neanderthal genetic variation in five key DHEAS biosynthesis genes (HSD3B2, CYP17A1, POR, CYB5A, SULT2A1). We also examine archaic introgression in these genes by comparing high-coverage Neanderthal genomes with globally diverse modern human sequences from the 1000 Genomes Project. We identify 29 Neanderthal-derived single nucleotide variants across these genes. Key steroidogenic genes in the biosynthesis pathway show no evidence of introgression, consistent with selection on pleiotropic regulators of steroidogenesis. In contrast, accessory genes carried introgressed Neanderthal haplotypes at moderate frequencies in non-African human populations, indicating Neanderthal variants are compatible with the human DHEAS synthesis pathway. All Neanderthal-specific variants were in non-coding regions, with three variants associated with reduced enzyme efficiency or DHEAS production in adults. Additionally, for all 29 positions, the modern human major allele is ancestral, and there is no evidence for a suite of novel adrenarche-extending variants. We conclude that the genetic foundation for extended adrenarche is shared between Homo sapiens and Neanderthals, and may have deeper hominin roots. Any phenotypic variation in adrenarche between modern humans and Neanderthals is more likely attributable to differential gene expression than to divergence in protein-coding sequences.
Moreau, C.; Morin, G.-P.; Bouchard, J.; Mathieu, J.; Duchesne, E.; Gagnon, C.; Girard, S. L.
Show abstract
Background: Myotonic dystrophy type 1 (DM1) is caused by a CTG repeat expansion in the DMPK gene and represents the most common adult-onset myopathy. Current molecular diagnostics rely on labor-intensive assays that limit accessibility and scalability. Haplotype-based approaches offer a promising alternative for detecting pathogenic expansions indirectly. Methods: We performed genome-wide genotyping in 226 genetically confirmed DM1 patients from the Saguenay-Lac-Saint-Jean founder population and reconstructed haplotypes surrounding the DMPK pathogenic repeat expansion. Based on these haplotypes, we performed a phylogenetic analysis that was further integrated with genealogical reconstruction from the BALSAC database to investigate the origin and transmission of DM1 haplotypes. To evaluate epidemiological utility, we implemented gene dropping simulations within the SLSJ extended genealogies (>80,000 starting individuals) to estimate DM1 incidence at birth. Results: A DM1-associated haplotype was identified in all patients (226/226), consistent with a single major ancestral origin in the SLSJ population. This complete concordance supports the robustness of haplotype-based approaches to infer carrier status without direct repeat sizing. Integrating phylogenetic analysis and genealogical data identified a single couple as the most likely entry point of DM1 in Quebec. Simulation-based estimates of incidence at birth exceeded observed prevalence, suggesting underdiagnosis in the region. Marked geographic heterogeneity in the SLSJ is also observed. Conclusions: Our results demonstrate that haplotype-based approaches can provide a reliable, cost-effective alternative to conventional pathogenic DM1 repeat carriers identification and familial screening strategies.
Roberts, M. C.; Jones, L. K.; Brown, A.; Carda-Auten, J.; Cuchel, M.; Hilton, A. R.; Khera, A.; Rothstein, M.; Soe, K.; Sullivan, A.; Tricou, E.; Vu, M. B.; Weintraub, W. S.; Ahmad, Z.
Show abstract
Objective: To identify patient- and clinician-reported barriers, facilitators, and design requirements for a centralized cascade-screening program for familial hypercholesterolemia (FH) in the United States. Methods: From June through November 2023, we conducted individual telephone interviews with 20 patients with FH and 10 clinicians recruited from UT Southwestern Medical Center, Parkland Health, the North Texas Veterans Affairs, and other clinical settings. Interview guides were informed by the Consolidated Framework for Implementation Research. Transcripts were coded in Dedoose using a piloted codebook, with discrepancies and emergent themes resolved through consensus. An advisory panel then helped translate interview findings into program design requirements and implementation strategies. Results: Five themes characterized barriers and facilitators to centralized cascade screening: (1) health-system access and fragmentation, including screening and treatment costs, transportation, and cross-system coordination; (2) privacy and trust, including concerns about genetic information and unsolicited outreach; (3) family relationships and practical burden, including competing demands, language barriers, limited contact, fear, and denial; (4) clinician capacity and workflow, including limited time, knowledge, and genetic-counseling capacity; and (5) communication and care continuity. Participants recommended proband pre-notification of relatives, culturally and linguistically responsive materials, secure data exchange, standardized scripts, flexible testing pathways, and centralized coordination. These findings informed a program model incorporating a secure pedigree platform, educational and communication resources, testing coordination, and linkage to follow-up care. Conclusions: Patients and clinicians identified multilevel determinants that a centralized FH cascade-screening program must address. The findings support specific design requirements but do not establish program feasibility or effectiveness, which require prospective evaluation.
Connors, P. D.; Guan, Y.; James, C. A.; Polaris, J.; Cantfil, B.; Campbell, C. A.
Show abstract
Purpose As genomics permeates all healthcare specialties, genetic counseling in conjunction with genetic testing is broadly recommended. Improving access to genetics professionals is crucial for Americans insured by Medicaid. The purpose of this study was to conduct a comprehensive review of Medicaid policies for genetic counseling performed by Certified Genetic Counselors (CGC(C)). Methods Fee-for-service Medicaid policies across 50 states and Washington DC were reviewed and coded. Four states with exemplary policies were identified, and CGC managers in two of these were surveyed regarding the real-world effects of these policies. Results As of 2024, 20 states (39%) had a published policy for genetic counseling with most (N=16, 80%) expressly covering genetic counseling in connection with any covered genetic test. Of the states without a policy, 12 (24%) mention genetic counseling in the context of scenario-specific policies, and 19 (37%) have no published policy. Twenty states explicitly cover CGC services, while 2 exclude CGCs as service providers. CGC managers in Indiana and Michigan confirmed the policies identified as exemplary successfully led to reimbursement of CGC services. Conclusion There is significant variability in Medicaid coverage for genetic counseling. Comprehensive policies are needed to support patient access to genetics professionals, including CGCs.
Yerukala Sathipati, S.; Scott, H.
Show abstract
Importance: Hereditary breast and ovarian cancer (HBOC) variant carriers benefit from risk-reducing interventions, but only if identified. The extent to which carriers are clinically recognized, and whether recognition is equitable across diverse populations, is poorly characterized in a single large U.S. cohort. Objective: To estimate P/LP HBOC carrier prevalence across genetic ancestry groups, quantify documented clinical genetic testing among carriers, and evaluate ancestry and socioeconomic disparities in testing. Design, Setting, and Participants: Cross-sectional analysis of the All of Us Research Program Controlled Tier (Curated Data Repository v8/C2024Q3R9), comprising participants with short-read whole genome sequencing and linked electronic health record (EHR) and survey data. Carriers were ascertained from research genomic data independent of clinical testing. Exposures: Genetically inferred ancestry (African [AFR], Admixed American [AMR], East Asian [EAS], European [EUR], Middle Eastern [MID], South Asian [SAS]); self-reported household income and educational attainment. Main Outcomes and Measures: (1) Carrier prevalence with Wilson 95% CIs; (2) documented clinical genetic testing (procedure codes) among carriers; (3) adjusted odds of documented testing among women, by ancestry, before and after socioeconomic adjustment, using multivariable logistic regression. Results: Among 414,830 participants, P/LP HBOC carrier prevalence was 1.42% (95% CI, 1.38-1.45) overall and similar across ancestry groups (AFR 1.24%, AMR 1.32%, EAS 1.19%, EUR 1.52%, MID 1.68%, SAS 1.33%; overlapping CIs). Among 250,071 women in the testing analysis, documented clinical genetic testing was rare: only 74 of 5,878 carriers overall (1.3%) and 59 of 3,572 European-ancestry carriers (1.7%) had a documented test, with counts below reportable thresholds in all other ancestry groups. African-ancestry women had lower adjusted odds of documented testing than European-ancestry women (Model 1 adjusted odds ratio [aOR], 0.32; 95% CI, 0.27-0.39), an association that attenuated but persisted after adjustment for income and education (Model 2 aOR, 0.48; 95% CI, 0.40-0.58; P < 0.001); Admixed American women also had reduced adjusted odds (aOR, 0.71; 95% CI, 0.61-0.84). Lower income and lower education were independently and dose-dependently associated with lower testing odds (income <$25,000 aOR, 0.46; high-school education aOR, 0.54). Conclusions and Relevance: High-risk HBOC variant carriers are present across all ancestry groups at similar frequencies, yet documented clinical genetic testing was disparate in the different ancestry groups. African-ancestry women experience a testing gap that is not fully explained by socioeconomic position, implicating structural barriers in access and referral. Population-level strategies that decouple carrier identification from current referral pathways may be required to close this gap.
Diaz-Papkovich, A.; Kuntzleman, A.; Davis, S. C.; Ramachandran, S.
Show abstract
Genetics research frequently intersects with ethnicity, nationality, and race, making it uniquely vulnerable to misrepresentation. Yet, 25 years after the initial sequencing of the human genome, there is little understanding of how human genetics research exists in the public-facing information ecosystem. We analyze 3,050,422 historical revisions from 6,738 Wikipedia pages about ethnicity, nationality, and race spanning 25 years. We find genetics terminology is present in 14.8% of these pages (55.5% in the top 1,000 pages) and in 67.8% of pages about nationalities, suggesting research is synthesized to present a biological element to ethnicity and nationality. We also find that 10.1% of 56,908 discussions from these pages contain genetics terminology. We further analyze responses from three popular chatbots queried about nationalities and find that they commonly reference both genetics and Wikipedia. Lastly, we analyze 133 pages from Grokipedia, an AI-generated encyclopedia, and find it mentions genetics more frequently than Wikipedia and hallucinates or misrepresents human genetics research.
Hoang, Q. P.; Le, T. X.; Doan, D. D.
Show abstract
Background. Polygenic scores (PRS) for coronary artery disease (CAD) are derived almost entirely from European-ancestry data. Their portability to Southeast Asian populations, including the Vietnamese, is largely uncharacterised and clinically consequential when scores are used with risk thresholds. Methods. We evaluated four independent European-derived CAD scores from the PGS Catalog (PGS000058, PGS000349, PGS002809, PGS004198; 70 - 5,723 variants) in 2,504 individuals from the 1000 Genomes Project, focusing on the Vietnamese Kinh (KHV) and Dai (CDX) samples. Per-individual scores were computed with PLINK2 and standardised. We assessed (i) the cross-ancestry distribution (calibration) and (ii) a clinically-relevant consequence: the proportion of each population flagged high genetic risk when the European top-20% threshold is applied (20% if perfectly calibrated). Results. For the primary score (PGS000058) the standardised PRS differed across super-populations (ANOVA F(4, 2499) = 121.1, p < 0.001); the Vietnamese Kinh mean was +0.47 SD above the European mean (Welch t = 7.77, p = 2.0 x 10^ -14). Applying the European top-20% high-risk threshold, the fraction of Vietnamese Kinh flagged ranged from 22.2% to 57.6% across the four scores, and of Dai from 21.5% to 43.0%, versus the intended 20%. Three of the four scores over-flagged Vietnamese (25-58%); the largest score (PGS004198) was approximately calibrated for East/Southeast Asians ([~]22%) but markedly over-flagged Africans (69.3%). Conclusions. European-derived CAD polygenic scores are inconsistently calibrated in Vietnamese and other Southeast Asian samples, and most substantially over-flag high genetic risk when a European threshold is applied. The magnitude and even the direction of miscalibration depend on the specific score, so no such score can be assumed transferable without local validation and recalibration. Distribution shift bounds, but does not by itself quantify, loss of predictive accuracy, which requires phenotyped data.
Horowitz, A. L.; Liebman, A. Z.; Liebman, S. W.
Show abstract
Founder mutations are variants that arose in a single ancestor and became enriched in a descendant population through a bottleneck and endogamy. Identification of pathogenic founder mutations has facilitated efficient targeted screening. More broadly, even without confirmed founder status, identifying pathogenic variants that are enriched within specific populations reveals population-specific disease burden. However, many such variants remain hidden in plain sight within existing datasets. To address this gap, we developed FIND (Founder candidates hidden IN Data), a web tool that identifies pathogenic, likely pathogenic, and predicted loss-of-function variants in gnomAD with frequencies >0.00008 in one ancestry group and at least tenfold higher than in all others (after zeroing populations with four or fewer observed alleles). Testing FIND on the genes FLNC, TMEM127, MYH7, and BRCA2 confirmed its utility and functionality by identifying nine well-known founder mutations and seven candidate founders. Candidates enriched in African American and admixed American populations were validated with the All of Us database, highlighting the utility of this approach for populations historically underrepresented in genetic studies. Source code is freely available at https://github.com/aacoder105/FIND under an MIT license, with a web interface at https://ethnic-variant-mutation-finder.onrender.com/.
Dong, R.; Wang, M.; Wang, G. T.; deWan, A. T.; Leal, S. M.
Show abstract
Motivation: Linkage disequilibrium score (LDSC) regression is a popular method to estimate heritability for complex traits using summary statistics and linkage disequilibrium (LD) reference panels, offering a practical alternative to methods requiring individual-level data. Despite its widespread use, LDSC regression can produce biased heritability estimates. The properties of LDSC regression were investigated using summary statistics from several large-scale Alzheimer's disease (AD) studies and a variety of LD reference panels. These heritability estimates were compared with those obtained from individual-level data. Results: When LDSC regression was applied to summary statistics obtained from meta-analysis, it led to an underestimation of heritability. This can occur if meta-analysis is used to combine studies of different ancestries leading to the caveat of the lack of an appropriate LD reference panel. Additionally meta-analyses often include studies with different phenotype definitions, that not only impacts heritability estimates but also makes them uninterpretable. Summary statistics generated from imputed variants, even those with high imputation accuracy, can lead to underestimation of heritability. For example, the heritability estimates for AD were reduced from 0.265 (se 0.148) to 0.160 (se 0.041) when imputed variants (INFO>0.9) were included compared to analyzing only genotype array variants. A decrease in heritability estimates was also observed when individual-level imputed variant data were analyzed using GCTA-GREML. Our findings highlight the caveats of estimating heritability using meta-analysis summary statistics or imputed data instead of genotyped or sequence data.
Sharma, J.; Maldonado, B.; Ungar, R. A.; Adimoelja, A.; Flores, J.; Gjorgjieva, T.; Jones, K.; Khan, A.; Xue, D.; Patel, R.; Caggiano, C.
Show abstract
Despite the importance of population descriptors in human genomics research, many scientists struggle to translate evolving ethical guidelines into their computational workflows. To characterize this gap between recommendations and implementation, we conducted a mixed-methods survey of early-career researchers to assess how they understand and implement the landmark 2023 NASEM report on the use of population descriptors in human genetics research. We show that while exposure to the report fosters ethical awareness, fundamental misconceptions about race and ancestry persist across academic disciplines, and trainees face structural bottlenecks, including legacy data constraints and a lack of technical confidence. To address this gap, we offer actionable, stakeholder-specific recommendations across the research lifecycle ranging from decision-support tools to "bring-your-own-data" workshops to leadership from academic journals, scientific societies, and trainee mentors. Ultimately, we argue that to promote scientific rigor and reduce bias in genetic discoveries, the scientific ecosystem must invest in the infrastructure necessary to empower the next generation of researchers.
Postma, J. K.; Haghshenas, S.; Bily, T. M.; Isovic, M.; White-Brown, A.; McConkey, H.; Kerkhof, J.; Rzasa, J.; Saleh, M.; Prasad, C.; Siu, V. M.; Carter, M. T.; Dyment, D. A.; Lazier, J.; Sawyer, S. L.; Moresco, A. A.; Jimena Diaz, M.; Abbate, S. L.; Campeau, P. M.; Innes, A. M.; Boycott, K. M.; Sadikovic, B.; Balci, T. B.
Show abstract
Background: Recurrent constellations of embryonic malformations (RCEMs) comprise multiple malformation conditions with largely unexplained etiologies and no established molecular biomarkers. A shared DNA methylation episignature was recently identified in VACTERL association and oculoauriculovertebral spectrum (OAVS). We sought to validate this episignature in an independent, deeply phenotyped cohort and evaluate its detection across related RCEMs. Methods: Genome-wide DNA methylation profiling was performed on peripheral blood from 38 participants with clinically diagnosed RCEMs, including VACTERL (n=21), partial VACTERL (n=3), OAVS (n=3), and other RCEM-related conditions (n=11). Results: The Episign V5 RCEM episignature demonstrated robust sensitivity for VACTERL (18/21, 85.7% positive), while the remaining three participants showed intermediate positivity. Of three participants with partial VACTERL, one with tracheoesophageal fistula demonstrated intermediate positivity, whereas the other two were negative. Episignature positivity was also identified in oculoauriculofrontonasal dysplasia (1/1, robust) and rhomboencephalosynapsis (1/2, robust) but was limited in OAVS (1/3, intermediate) and absent in frontonasal dysplasia (0/4). Conclusions: Independent validation establishes the Episign V5 RCEM episignature as a reproducible molecular biomarker for VACTERL, a condition that remains a diagnosis of exclusion. Variable detection across related malformation conditions suggests etiologic heterogeneity, whereas overlap among selected phenotypes supports epigenomic convergence across the RCEM spectrum.